Laboratory Investigation
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Laboratory Investigation's content profile, based on 13 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Burley, A.; Silveira, T.; James, N.; Salto-Tellez, M.; Wilkins, A. C.
Show abstract
Background: Single cell RNA sequencing provides a wealth of information to explore the complexities of the tumour microenvironment, but crucially the spatial topology of the tumour is lost and studying cellular interactions is limited. Spatial transcriptomics aims to address this however the technique remains cost prohibitive for the generation of data from meaningfully-sized clinical cohorts. In contrast, spatial proteomic profiling with multiplex immunofluorescence, preserves spatial interactions, is relatively cost accessible, and is scalable for large clinical cohorts to address powerful translational questions. Whilst multiplex approaches have advanced in recent years, we note that cancer-associated fibroblasts (CAFs) have been explored in less detail, potentially due to difficulties associated with CAF heterogeneity and the diversity of markers used to define them. Methods: We designed, optimised, and validated a multiplex immunofluorescence panel that combines four frequently used CAF markers; alpha smooth muscle actin (aSMA), fibroblast activation protein (FAP), podoplanin (PDPN) and platelet-derived growth factor receptor alpha (PDGFRa) with CD8 and pan-cytokeratin. Here we share our methodology and the practical considerations taken to inform the final panel design. We also highlight the benefits of robust optimisation experiments.
Fernandes, G. M. d. M.; Wang, W.; Parwani, A.; Ahmadian, S. S.; Alves, M. J.; Philips, J. J.; Otero, J. J.
Show abstract
The reproducibility of immunohistochemistry in tumor tissue analysis across reference labs remains a persistent challenge. We tested the extent to which an intra-slide calibration technology mitigated discprepencies in inter-laboratory assays of p53 immunohistochemical (IHC) reactions in brain biopsies of glioblastoma (GB), IDH-wildtype. Intra-slide calibration technologies apply a 0-100% concentration scale incorporating primary surrogate and secondary antibodies to generate a standardized curve for DAB precipitation. IHC from GB samples was performed independently by pathology departments from two different hospital laboratories and were digitalized at 40x magnification using Aperio Image Scope software. Feature extraction, including intensity and texture parameters was performed using the EBImage package in R, followed by UMAP dimensionality reduction and DBSCAN clustering analysis. Our results show significant differences in intensity and texture clustering patterns between laboratory tissue samples and intra-slide calibration technology ruler caused by the different laboratories. Intra-slide calibration technology coupled with polynomial regression analysis improved ~90% the data harmonization. Our findings demonstrate a key role for computational pathology using intra-slide calibration technology to enable intra-laboratory consistency and inter-laboratory reproducibility. These advances strengthen the reproducibility of diagnostic assessments and support more objective, data-driven decision-making in neuro-oncology.
Guedes, J.; Sliwa-Gonzalez, A.; Szadai, L.; Geiger, P.; Woldmar, N.; Reyes, M. A.; Bastida, R. A.; Coto, D. L. F.; Oskolas, H.; Marko-Varga, M.; Schultz, L.; Appelqvist, R.; Wieslander, E.; Malm, J.; Marko-Varga, G.; Gil, J.
Show abstract
Melanoma incidence continues to rise globally, with formalin-fixed paraffin-embedded (FFPE) tissue archives representing an invaluable resource for large-scale retrospective proteomic studies. However, inconsistent deparaffinization remains a critical pre-analytical bottleneck limiting protein yield, reproducibility, and downstream data quality. In this study, we developed and validated a fully automated FFPE deparaffinization workflow using the Fluent(R) 780 liquid handling workstation (Tecan (C)) and evaluated its performance against a conventional manual protocol in a cohort of 54 patients with primary cutaneous melanoma, predominantly at early AJCC 8th edition stage I-II. The automated workflow achieved superior protein identification (6,146 {+/-} 860 vs. 4,941 {+/-} 1,091 proteins; p < 0.0001) with lower technical variability, while maintaining highly comparable global proteomic profiles as confirmed by principal component analysis and hierarchical clustering. A total of 8,305 proteins (96.1%) were identified by both methods, supporting the reproducibility and equivalence of the automated approach. Patients were stratified by the presence (N=21) or absence (N=33) of histological regression in the primary tumor. Proteomic comparison revealed 97 upregulated and 226 downregulated proteins in regressing melanomas, with pathway enrichment analysis demonstrating elevated mitochondrial and translational activity alongside reduced innate immune and complement pathway activation in the regression group. No statistically significant differences in overall, disease-free, or progression-free survival were observed between groups, consistent with the early-stage composition of the cohort. Digital pathology validated tissue morphology preservation across processing conditions. These findings support the integration of automated FFPE processing with proteomic and digital pathology workflows as a scalable platform for precision melanoma research. TOC Figure O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=133 SRC="FIGDIR/small/744404v1_ufig1.gif" ALT="Figure 1"> View larger version (49K): org.highwire.dtl.DTLVardef@1d51629org.highwire.dtl.DTLVardef@a1f126org.highwire.dtl.DTLVardef@1df1b0aorg.highwire.dtl.DTLVardef@686f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Yang, X.; Marlin, M. C.; Celia, A. I.; Lee, C.-Y.; Cammarata-Mouchtouris, A.; Stephens, T.; Haddad, M.; Bradshaw, L.; Saksena, D.; Buyon, J.; Izmirly, P. M.; Putterman, C.; Kamen, D.; Petri, M.; Accelerating Medicines Partnership: RA/SLE Network, ; James, J. A.; Guthridge, J. M.; Fava, A.; Rosenberg, A. Z.
Show abstract
BackgroundTraditional immunohistochemistry (IHC) with chromogen detection has limited multiplex capacity, detecting at most 4 protein markers per tissue section simultaneously, thereby restricting comprehensive spatial analysis of valuable human biopsies. We developed and validated a robust serial IHC (sIHC) staining method to detect multiple antigens on a single kidney biopsy slide, maximizing data yield for diagnosing and studying complex kidney diseases. MethodsFormalin-fixed, paraffin-embedded kidney biopsy sections were subjected to repeated IHC/imaging cycles with antibody removal using an optimized sodium dodecyl sulfate-glycerol buffer stripping protocol. Images were then co-registered, and analysis was performed using a variety of methodologies, including color deconvolution, cell segmentation, and spatial clustering. ResultsThis optimized sIHC method successfully detected up to 20 antigens on a single slide. Combining image analysis and artificial intelligence software, for example with HALO (Indica Labs), the assay assembles high-dimensional images and enables quantitative histology and single-cell spatial analysis. Using this advanced method, we were able to identify rare cell populations, such as double-negative T cells, that are challenging to detect conventionally. ConclusionWe have developed a validated, high-capacity sIHC protocol that uses standard IHC procedures with commercially available, clinically validated off-the-shelf antibodies. This method is a valuable, cost-effective tool for obtaining extensive, high-dimensional single-cell-resolved spatial data from limited pathology samples, such as a human kidney biopsy.
Grion, G.; Hussain, R.; Colella, F. E.; Roufail, K.; Uccella, S.; Frapolli, R.; Matteo, C.; Mintemur, O.; Pennati, F.; Renne, S. L.
Show abstract
Quantifying vascular architecture in histological whole slide images is needed to study tissue organisation, tumour microenvironment biology, and diseaseassociated vascular remodelling. However, vessel analysis in routine immunohistochemistry remains challenging. Available workflows are often manual, require programming expertise, or lack direct integration with digital pathology platforms. We developed VeSpA (Vessel Spatial Analysis), an open-source pipeline and QuPath extension for automated vessel segmentation and morphometric quantification in CD31-stained whole slide images. VeSpA combines configurable signal extraction, using CMYK Yellow channel extraction by default and optional DAB stain deconvolution for H-DAB images, with automatic or percentile-based thresholding, morphological refinement, contour filtering, and lumen filling to generate vessel masks from standard DAB-stained sections. The QuPath extension includes a graphical interface for selecting annotations, TMA cores, or whole images, configuring segmentation parameters, running the Python backend, and importing vessel objects directly into the QuPath hierarchy. For each detected vessel, VeSpA extracts area, major axis length, minor axis length, eccentricity, centroid, and orientation, while also appending summary measurements to parent annotations and TMA cores. Validation against independent pathologist annotations showed that VeSpA achieved segmentation performance close to inter-rater agreement and outperformed yellow channel prompt-based SAM and zero-shot YOLOv8-seg on overlap-based metrics in the tested dataset. VeSpA integrates vessel segmentation, morphometric feature extraction, and QuPath-based visualisation into a single reproducible workflow for vascular quantification in computational pathology and spatial analysis of histological tissue architecture.
Rounds, C. C.; Ravi, D.; Huang, G.; Mengesha, B.; Tran, S.; Garcia, A.; Rueb, N.; Chang, Y. H.; Park, B. S.; Wong, M. H.; Gibbs, S. L.
Show abstract
SignificanceRare-cell identification in fluorescence microscopy remains challenging because targets are sparse and background varies between specimens. Combining specimen-specific fluorescence enrichment with image classification may enable efficient and more specific automated detection of rare cells. AimWe developed a two-stage framework to identify and quantify candidate rare circulating hybrid neoplastic cells (CHCs, ECAD+/CD45+) in peripheral blood mononuclear cell (PBMC) preparations from tumor-bearing and tumor-naive mice. ApproachPBMCs from 28 mice were imaged by multichannel fluorescence microscopy. Matched unstained samples established animal-specific ECAD and CD45 background distributions for candidate cell enrichment. Blinded multi-annotator consensus labels were used to train a convolutional neural network (CNN) from DAPI, ECAD, and CD45 image crops. Generalization was evaluated by leave-one-animal-out validation across 10 random initializations. Final classification used a 10-model ensemble, and rare-cell burden was compared between groups using negative-binomial regression with total segmented-cell count as an exposure. ResultsOf the 1,065,512 segmented cells, enrichment retained 10,176 candidates (0.96%), reducing the search space by >99%. Four of five evaluable tumor-bearing animals showed reproducible held-out discrimination, with median quantified area under the receiver operator characteristic curve (AUROCs) of 0.918-0.951; one animal was a reproducible outlier (median AUROC, 0.338). Ensemble deployment identified 157.94 positive-consensus cells per 50,000 segmented cells in tumor-bearing animals versus 49.55 in controls. The estimated rare-cell rate was 3.15-fold higher in tumor-bearing animals (95% CI, 0.91-10.99; two-sided p=0.071; prespecified one-sided p=0.036). ConclusionsSpecimen-specific fluorescence enrichment combined with supervised image classification reduced the cellular search space and enabled automated quantification of a rare CHC (ECAD+/CD45+) phenotypes. Cross-animal validation also identified specimen-specific generalization failure, highlighting the importance of biological-specimen-level validation.
Castillo, S. P.; Gautam, T.; Pinao Gonzales, K. B.; Salvatierra, M. E.; Serrano, A.; Ercan, C.; Rodriguez, B. L.; Acosta, P.; Chen, P.; Shokrollahi, Y.; Lau, A.; Kwong, L. N.; Huse, J. T.; Pan, X.; Patient Mosaic Team, ; Solis Soto, L. M.; Yuan, Y.
Show abstract
Selection of regions of interest (ROIs) is often a crucial step in spatial molecular profiling and many pathology tasks, with substantial implications for research reproducibility and biological interpretability. To provide a reproducible and adaptive framework for AI-guided ROI selection, we developed a modular generalist-specialist solution across spatial profiling platforms. In a cohort comprising 55 tumor types from 160 tissue donors profiled using NanoString Digital Spatial Profiling and multiplex immunofluorescence, we first established a protein-profiling reference atlas capturing compartment-specific immune, checkpoint, stromal, and proliferation patterns. We then developed an AI Specialist Task-Oriented Model for ROI Selection (ASTROS) and tested comprehensive benchmarks considering specialist-only (ASTROS), generalist-only (PLIP/GFM), and hybrid generalist-specialist strategies, showing that the latter provides a balanced tradeoff across slide-level signal preservation, pathologist-reference concordance, within-slide placement consistency, and large-slide computational efficiency. We further demonstrated the feasibility of virtual staining for ROI preview and modular ROI placement for other spatial omics technologies, Visium and Visium HD workflows. Together, these results support our proposed framework to enable ROI selection responding to unmet needs for reducing inter-rater variability, reproducibility, and versatility in spatial profiling experiments.
Kuempers, C.; Stein, K.; Nitschkowski, D.; Jagomast, T.; Heidel, C.; Kirfel, J.; Droemann, D.; Bohnet, S.; Schweigert, M.; Reck, M.; Olchers, T.; von Weihe, S.; Ammerpohl, O.; Goldmann, T.
Show abstract
Non-small cell lung cancer (NSCLC) is the most common form of lung cancer accounting for most cancer-related deaths worldwide. Despite substantial recent advances in targeted therapies and immunotherapy, the prognosis for advanced-stage disease remains comparably poor, which is why the identification of novel molecular biomarkers as well as therapeutic targets influencing tumor development, progression, and metastasis remain important. This study focusses on SERPINB13, a serine-protease inhibitor expressed in selected tissues that is dysregulated in several tumor entities. However, its role in NSCLC still remains largely unclear. We analyzed SERPINB13 transcription in a cohort of non-small cell lung cancer (NSCLC) cases including both lung squamous cell carcinoma (LUSC) and lung adenocarcinoma (LUAD) by transcriptome profiling. Epigenetic modifications were assessed via Methylation BeadChips. Additionally, SERPINB13 protein expression was assessed by immunohistochemistry (IHC) in an independent cohort of NSCLC comprising 126 LUSC patients. Correlation analyses were performed to associate SERPINB13 expression with key clinico-pathological parameters, including overall survival and extent of tumor-infiltrating immune cells. To functionally investigate the regulatory influence of peripheral blood mononuclear cells (PBMCs) on SERPINB13 expression in LUSC tumor cells in vitro, we utilized the SERPINB13-expressing LUSC cell line LUDLU-1. Here, gene transcription was analyzed by quantitative real-time PCR (RT-qPCR), confirmed by Western blot on the protein level. Transcriptome analysis revealed a significant upregulation of SERPINB13 in lung squamous cell carcinoma (LUSC) compared to lung adenocarcinoma (LUAD), highlighting a subtype-specific expression pattern. This differential expression was further associated with a distinct epigenetic DNA methylation signature at the SERPINB13 loci in LUSC, suggesting transcriptional regulation via hypomethylation. IHC analysis demonstrated that high SERPINB13 protein expression is significantly associated with prolonged overall survival in LUSC. Notably, SERPINB13 expression was enriched in immune-inflamed ("hot") tumors, characterized by elevated infiltrating lymphocytes and immune activation. Mechanistically, co-culture experiments with PBMCs induced SERPINB13 expression in a LUSC cell line in a dose- and time-dependent manner in the absence of direct cell contact. This suggests that soluble factors secreted by immune cells might play a key role in regulating SERPINB13 expression in the tumor microenvironment. Taken together, SERPINB13 is a novel prognostic indicator in LUSC that is modulated by Immune cells. Further studies are necessary to decipher the crosstalk of Immune cells on the Serpin B13 expressing tumor cells in depth, with regard to a possible interventional strategy. immunomodulatory potential strategies and personalized therapeutic approaches in NSCLC.
Wouters, J.; Bertorello, J.; Gaillard, M.; Simon, B.; Gastineau, S.; Roehrig, A.; Dupont-Roc, M.; Amblard, E.; Pupo, A.; Yu, H.; Blay, J.-Y.; Guerin, C.; Nebot Bral, L.; Vincent Salomon, A.; Verlingue, L.; Xylina, E.; Cabel, L.; Ross, J.; Miller, V.; Letouze, E.; Vallot, C.
Show abstract
Tumor cellular composition--including malignant cell states, immune populations, and stromal populations--is increasingly recognized as a determinant of therapeutic response and resistance to anti-cancer agents, yet comprehensive cellular profiling remains largely confined to research settings. Here, we present a clinically compatible sample-to-report workflow for tumor composition profiling from routine formalin-fixed paraffin-embedded (FFPE) clinical specimens. By combining low-input single-nucleus RNA sequencing with foundation model- based automated cell annotation, this workflow enables prospective sample-by-sample analysis without dedicated research material or cohort-based processing. Across 116 clinical specimens representing six cancer types, we generated reproducible measurements of cellular composition and cell-type-specific gene expression, demonstrated high technical reproducibility, and showed concordance with pathological assessment of immune infiltration. The workflow was similarly applicable to archival FFPE material and ultra-low-input biopsy specimens. Together, these findings establish a practical framework for routine single-cell profiling from standard pathology specimens and open the perspective of prospective evaluation of cellular composition as a clinical biomarker in precision oncology.
Cervantes-Rivera, R.; Figueroa Ortiz, S. J.; Romero Rosas, A. Z.; Sanchez Orozco, A.; Herrera-Vargas, M. A.; Melendez-Herrera, E.; Lopez-Rodriguez, M.; Ochoa-Zarzosa, A.; Lopez-Meza, J. E.
Show abstract
Three-dimensional (3D) spheroid models have become essential in cancer biology, drug screening, and tissue engineering. However, their small size, fragile structure, and tendency to disintegrate during routine histoprocessing present persistent technical challenges. Conventional paraffin embedding often results in tissue fragmentation, loss of spatial orientation, and poor section quality, whereas cryosectioning often compromises cellular morphology. Here, we present a robust, cost-effective protocol for preserving and sectioning fragile 3D spheroids, resulting in high-quality histological sections with intact architecture and excellent cellular detail. The method involves optimized handling and embedding procedures that stabilize spheroids during standard formalin fixation, paraffin infiltration, and microtomy, eliminating mechanical distortion and preserving spherical integrity for consistent sectioning. We demonstrate successful application across different cell line spheroids, with subsequent compatibility with hematoxylin and eosin (H&E) staining protocols. Compared to conventional methods, our approach significantly reduces sample loss, improves inter-section reproducibility, and preserves fine structural features such as necrotic cores, proliferative zones, and extracellular matrix components. This protocol provides a reliable, accessible solution for routine histological analysis of fragile 3D spheroids, facilitating more accurate morphological and molecular assessment in translational research settings. Key featuresO_LIMaintains spheroid integrity: Prevents mechanical distortion, fragmentation, and loss of spatial orientation during processing. C_LIO_LISignificantly reduces sample loss: Decreases failure rate compared to traditional methods, conserving valuable samples. C_LIO_LIBroad spheroid compatibility: Works effectively with primary tumor-derived, stem cell-derived, and co-culture spheroid models. C_LIO_LIEnables high-quality sectioning and staining: Delivers consistent, reproducible sections that are fully compatible with H&E, IHC, and IF. C_LI Graphical overview O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=140 SRC="FIGDIR/small/743094v1_ufig1.gif" ALT="Figure 1"> View larger version (44K): org.highwire.dtl.DTLVardef@1670c4org.highwire.dtl.DTLVardef@145810aorg.highwire.dtl.DTLVardef@1accb1org.highwire.dtl.DTLVardef@17481c0_HPS_FORMAT_FIGEXP M_FIG C_FIG
Calapaqui Teran, A. K.; Gonzalez Bernad, A. A.; Cobo Cano, M.; Sanchez Magdaleno, L.; Marcos Gonzalez, S.; Delgado Bolton, R. C.; Moustafa Calvo, J.; Gomez Roman, J. J.; Lara, L.
Show abstract
We present PRECISE (PRostate Expert-annotated Contiguous IHC-H\&E Serial sEctions), a hybrid histopathology dataset of paired hematoxylin and eosin (H\&E) and immunohistochemistry (IHC) whole-slide images (WSIs), comprising 37 prostate core needle biopsies from 25 patients, each with matched H\&E and CKAPM+racemase staining. To the best of our knowledge, this is the first publicly available dataset offering spatially harmonized, pixel-level expert annotations across both staining modalities in prostate biopsy WSIs - directly mirroring the two-stage (H\&E-then-IHC) clinical diagnostic workflow used to resolve morphological uncertainty, restricted to cases in which that workflow reached diagnostic consensus. The dataset contains 24,387 annotations spanning seven diagnostically critical classes: malignant glands, benign glands, stromal tissue, intraductal carcinoma (IDC-P), high-grade prostatic intraepithelial neoplasia (HGPIN), atypical intraductal proliferation (AIP), and tissue artifacts. Unlike existing resources, which focus on binary tumor classification or lack IHC pairing, this dataset captures the full morphological spectrum encountered in routine prostate pathology, including rare precursor lesions and confounding entities underrepresented in current benchmarks. Annotations were validated through a structured three-stage consensus by two expert uropathologists, with IHC serving as biological ground truth for boundary definition. PRECISE is designed as a robust benchmark for multimodal semantic segmentation and self-supervised learning, and is openly released to promote reproducible research and accelerate AI-assisted diagnosis in prostate cancer.
Maurer, J.; Suzuki-Horiuchi, Y.; Duong, B.; Ramirez, M. V.; Chen, A.; Prouty, S. M.; Milman, T.; Lee, V.; Cheng, Y.
Show abstract
Introduction Conjunctival melanoma (CM) is a rare cancer with a potentially high recurrence rate. The mechanics of its progression, its relationship with neighboring tissues, and its molecular characteristics are largely unknown. Diagnosis currently requires a biopsy and the time and expertise of a pathologist. Methods Archived human biopsies containing CM were submitted to Xenium spatial transcriptomic analysis. Regions were graded by disease progression through histopathology. Differential expression (DE) and composition analysis were performed across disease states. Results From three patients, 12 formalin-fixed paraffin-embedded (FFPE) tissue specimens were recovered. Composition analysis showed that melanoma depletes fibroblast and epithelial cells while melanocytes proliferate. DE signatures specific to each state show a clear pattern of progression from inflammation, to cellular restructuring, and then to tumor progression and malignancy. Conclusion Spatial transcriptomics allows single-cell transcriptomics techniques to compare spatially relevant annotations that are difficult to separate by library. This study proposes disease progression biomarker candidates that may elucidate the mechanics of CM progression and function as objective diagnostic and prognostic tools in the future.
Kang, Y.-J.; Jun, S.-Y.; Kim, S.
Show abstract
Background. Breast cancer treatment depends on histopathological features, such as grade and receptor-defined subtype; however, specialist pathologist access is constrained when the workforce is limited. Commercial multimodal large language models (MLLMs) accept hematoxylin and eosin (H&E) image tiles through paid interfaces without local hardware or fine-tuning. However, prior pathology evaluations addressed only coarse tasks. Whether they reach treatment-determining accuracy and whether vendors agree remain unclear. Methods. We aimed to evaluate three vendor-designated flagship MLLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) in 427 invasive breast cancer cases. Each case went to all three with identical H&E tiles and prompts, and the subtype was inferred in the second call. The reference was an institutional sign-out report of an immunohistochemistry-derived subtype. We calculated the concordance, sensitivity, specificity, Cohen's kappa, and pairwise McNemar and Bowker tests. Findings. Claude ranked highest by raw histologic-type concordance but lowest by kappa, classifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. The models anchored the Nottingham grade to three modal grades. None of the models reliably identified human epidermal growth factor receptor 2-positive disease. The failure direction was vendor-specific: Claude and GPT-5.5 were under-detected, whereas Gemini was over-called. Twelve prompt variants (4,056 calls) did not recover sensitivity. Interpretation. No current commercial MLLM reaches deployment-ready accuracy for any treatment-determining feature of breast pathology. As each vendor fails in its own fixed direction, changing vendors alters the type of error rather than removing it; therefore, the value of these models is assistive rather than autonomous. At USD 0.20-0.50 per case, they may serve as supervised draft generators that leave the diagnosis with the pathologist.
NYMAN, P.; Tampu, I. E.; Shamikh, A.; Prochazka, G.; Blystad, i.; Basmaci, E.; Diaz de Stahl, T.; Augustsson, P.; Zielinska-Chomej, K.; Cao, D.; von Salome, J.; Ardalan, A.; Somarajan, P. R.; Ljungman, G.; Lundberg, P.; Sandgren, J.; Haj-Hosseini, N.
Show abstract
Refined detection methods, more detailed tumor characterization, and adequate distinction between different pediatric tumor subtypes are necessary to improve diagnosis and treatment, enable precision medicine, and advance patient prognosis. However, the application of computational approaches to pediatric brain tumors remains limited, largely due to the lack of accessible datasets. To address part of this gap, we provide whole slide images (WSIs) of hematoxylin and eosin (H&E)-stained tissue sections from all pediatric central nervous system (CNS) samples collected in Sweden between 2013 and 2023. These data represent a population-based national cohort encompassing all six pediatric oncology centers in Sweden and are available through the Swedish Childhood Tumor Biobank (BTB). The dataset includes 1,446 WSIs of sufficient image quality with confirmed CNS tumor diagnoses, derived from 537 unique subjects (562 cases). In addition, diagnosticrelevant clinical information is included. Corresponding whole-genome sequencing (WGS), wholetranscriptome sequencing (WTS), and methylation array data are available for most tumor samples through separate resources. This H&E dataset has been specifically curated to support artificial intelligence-based analyses, while also serving broader applications in medical research and education. When combined with matched molecular data, it provides a valuable resource for advancing multimodal and precision diagnostic approaches in the pediatric population. Refined detection methods, more detailed tumor mapping and adequate distinction between different subtypes of pediatric tumors are necessary to improve treatment, enable precision medicine and improve patient prognosis. Application of computational algorithms for pediatric brain tumors is very limited mainly due to the unavailability of pediatric histology brain tumor data sets. To enable the development of AI models comprehensive datasets covering a wide range of pediatric brain tumors are needed.
Hofstraat-Boersma, R.; du Long, R.; Buzzanca, G.; Abiola, A. A.; Albadri, S.; Ali, Z.; Altaleb, A.; Angioi, A.; Banu, S. G.; Barry, M.; Bhalodia, A. R.; Bianco, P.; Broecker, V.; Buelow, R.; Chauveau, B.; Chen, G.; Cheunsuchon, B.; Crisi, G. M.; Daneshvar, S.; Dendooven, A.; Dokouhaki, P.; Drachenberg, C. B.; Farris, A. B.; Ferlicot, S.; Florquin, S.; Fontana, F.; Gibier, J.-B.; Gibson, I. W.; Gujarathi, S.; Hendricks, A. R.; Husain, S.; Islam, J.; Ismail, W.; Jagannathan, G.; Klager, J.; Kozakowski, N.; Krizova, A.; Kurien, A. A.; Kwon, B.; L'Imperio, V.; Ledesma, F. L.; Low, J. P.; Martin, J
Show abstract
Background Diagnostic interpretation of kidney allograft biopsies using the Banff classification remains variable, but the determinants of this variability are not fully defined. We performed a global, fully digital multi-reader study to identify the principal drivers of disagreement in Banff-based assessment. Methods Thirty six kidney transplant biopsies were independently scored by 67 renal pathologists on a standardized digital platform. Readers assessed Banff lesions on hematoxylin and eosin, periodic acid Schiff, and Jones' silver stains; final diagnostic categories were assigned using prespecified Banff-based decision rules. Interobserver agreement was quantified with Gwet's agreement coefficient (AC) statistics. Determinants of diagnostic agreement were evaluated) using pairwise mixed-effects logistic regression, and reader similarity was examined by principal component analysis (PCA) with post hoc molecular annotation. Results Agreement for final diagnostic categories was moderate (Gwet's AC1, 0.55; 95% CI, 0.47 - 0.63). Lesion-level agreement varied substantially, with lowest agreement for selected threshold-dependent inflammatory or semi-quantitative lesions, including interstitial inflammation in areas of IFTA, peritubular capillaritis and arteriolar hyalinosis. Diagnostic concordance differed markedly across biopsies, indicating strong case-level heterogeneity. In pairwise models, differences in active inflammatory and vascular lesion scoring were the strongest correlates of diagnostic disagreement; reader experience and geography contributed minimally. Principal component analysis showed reader variation was organized along two dominant axes: a rejection-calling threshold axis linked mainly to tubulointerstitial inflammatory injury, and a T cell-mediated (TCMR/TI) and antibody-mediated/microvascular (AMR/MVI) inflammation-oriented phenotypic classification axis. Conclusion Interobserver variation in Banff-based kidney transplant biopsy assessment is structured rather than random and driven mainly by how readers threshold and integrate key inflammatory lesion compartments rather than experience or geographic location.
Kuempers, C.; Roettger, H.; Jagomast, T.; Emken, L.; Heidel, C.; Paulsen, F.-O.; Tuecking, T.; Kirfel, J.; Droemann, D.; Bohnet, S.; Schweigert, M.; Reck, M.; Olchers, T.; von Weihe, S.; Meidl, V.; Nitschkowski, D.; Goldmann, T.
Show abstract
P2Y12 receptor (P2RY12), mainly expressed on platelets, is known for its central role in hemostasis. P2RY12 activation is also involved in cancer development through platelet adhesion to cancer cells supporting immune-evasion, promoting tumor angiogenesis and metastasis, among others. P2RY12 is known as an actionable target, and P2RY12 antagonists are in clinical use for cardiovascular diseases. However, very little data are available regarding the protein expression of P2RY12 in lung carcinomas. We performed immunohistochemical staining for P2RY12 in a cohort of non-small cell lung cancer (NSCLC) samples comprising 320 adenocarcinomas (LUAD) and 158 squamous cell carcinomas (LUSC). Results were evaluated using a dual approach combining microscopic assessment and digital image analysis (QuPath). Results were correlated with clinical-pathological data. We found significantly higher P2RY12 protein expression in LUSC compared to LUAD (p<0.001) via eyeballing (absent/low expression in 21.7% (34/158) and moderate/high expression in 78.3% (124/158) of LUSC cases versus absent/low expression in 98.4% (315/320) and moderate/high expression in 1.6% (5/320) of LUAD cases). Digital analysis yielded similar results. High P2RY12 expression was associated with a significantly better 5-year overall survival rate for the entire cohort (p=0.0048) as well as for the LUAD (p=0.015) and LUSC (p=0.05) subgroups. Furthermore, P2RY12 showed excellent discriminatory performance for classifying carcinomas as LUAD or LUSC, with an AUC of 0.916 in ROC-analysis. High P2RY12 expression is linked to a better prognosis and might serve as a promising novel prognostic biomarker for NSCLC. Its assessment could be implemented in future routine diagnostic workup. At the same time, the data suggest that P2RY12 could also serve as a diagnostic marker for LUSC.
Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.
Show abstract
Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.
Yuan, T.; Bai, Y.; Song, L.; Zhang, A.; Liu, Y.; Cao, Y.
Show abstract
Cell-free DNA (cfDNA) methylation profiling holds great promise for non-invasive cancer detection, yet accurate methylome analysis is compromised by DNA damage inherent to cfDNA. Standard library preparation workflows involve an end-repair step during which DNA polymerases can initiate synthesis from single-strand breaks (nicks), replacing endogenous methylated nucleotides with unmethylated nucleotides in the 3' direction and systematically erasing methylation information. This artifact is distinct from the terminal jagged-end effect and disproportionately affects cfDNA and FFPE DNA, which harbor abundant nicks. Here, we developed cf-Cabernet, which builds upon the Cabernet framework (Cao et al., 2023) -- an enzymatic methylation sequencing method featuring carrier DNA-assisted sample recovery and amplification-friendly post-conversion processing -- with the addition of a Taq DNA ligase-mediated nick repair step prior to end repair. Using matched cfDNA samples, we compared cf-Cabernet against standard EM-seq and WGBS. cf-Cabernet and EM-seq both substantially outperformed WGBS in alignment rate. Critically, while standard EM-seq exhibited globally reduced methylation levels compared to WGBS, cf-Cabernet yielded methylation levels concordant with WGBS. M-bias analysis revealed that EM-seq libraries showed persistently depressed methylation across the entire read length, whereas cf-Cabernet methylation recovered to the WGBS baseline beyond the terminal [~]40 bp jagged-end region. Nick-induced methylation erasure during end repair is a significant but previously underappreciated source of error in cfDNA methylation sequencing. cf-Cabernet effectively mitigates this artifact through pre-emptive nick ligation, enabling accurate methylome profiling from damaged DNA templates. This method is broadly applicable to cfDNA, FFPE DNA, and other clinical specimens where DNA integrity is compromised, providing a robust foundation for methylation-based liquid biopsy applications.
Hulahan, T. S.; Spruill, L.; Gerding, B. E.; Wang, M.; Macdonald, J. K.; Taylor, H. B.; Wallace, E.; Strand, S. H.; Mehta, A. S.; Ford, M. E.; Nakshatri, H.; Marks, J. R.; Angelo, M.; Colditz, G. A.; Hwang, E. S.; Drake, R. R.; West, R. B.; M Angel, P. M.
Show abstract
BackgroundDuctal carcinoma in situ (DCIS) is a noninvasive breast lesion with variable risk of progression to invasive breast cancer (IBC). Current transcription and cell marker investigations suggest ECM decreases in later events but are limited in details of ECM proteomic composition, including post-translational modifications. We investigated whether the extracellular matrix (ECM) proteome alters with later breast events of DCIS or IBC. MethodsECM-targeted mass spectrometry imaging and liquid chromatography-tandem mass spectrometry (LC-MS/MS) were applied to ten tissue microarrays from the Resource of Archival Human Breast Tissue cohort (RAHBT). Primary DCIS specimens (n=136) were analyzed in relation to later events of DCIS (n=40) or IBC(n=30), with a mean follow-up of 192.1 months 95% CI [179.1,205.1]. Statistical modeling, survival analyses, and exploratory machine learning approaches were used to identify ECM peptide signatures associated with later events. ResultsDistinct ECM peptide profiles were associated with later events of DCIS or IBC. Fifteen peptides derived from fibrillar collagens (COL1A1, COL1A2, COL3A1) and elastin, showed significantly reduced abundance in patients who developed IBC. Lower expression of specific collagen peptides associated with overall 19.9% 95% CI [17.92, 21.81] decreased disease-free survival for IBC. Lower expression of these peptides was significantly associated with reduced disease-free survival (age-adjusted hazard ratio [HR] = 2.45, 95% CI: 2.33-2.57; P < 0.05). Patient-matched samples of primary DCIS, later DCIS, and later invasive breast cancer further demonstrated reduction in ECM peptide detection. Exploratory predictive modeling from patient-matched samples achieved high performance (AUROC >0.98, accuracy >93%) in distinguishing primary from later events. Following prior work in the RAHBT cohort, reduction of certain collagen peptides was also observed in primary DCIS samples from higher risk patient groups. ConclusionsECM proteomic remodeling, particularly decreases of specific collagen domains, is strongly associated with later events of DCIS and IBC. These findings highlight ECM proteome as a critical regulator of breast cancer emergence with potential as a prognosticator of risk stratification to guide clinical management of DCIS.
Wang, E.; Grenier, K.; Savadjiev, P.; Poenaru, D. D.
Show abstract
Background. Definitive diagnosis of Hirschsprung's disease (HD) requires pathological identification of enteric ganglion cells. This process is time-consuming and subject to inter-observer variability. Artificial intelligence (AI) tools have the potential to standardize and accelerate this workflow, but no study has determined which AI approach best serves intraoperative HD pathology diagnostics. Method. This study compared the U-Net and You Only Look Once version 26 (YOLO26) frameworks for ganglion cell detection using a single-centre retrospective dataset of 54 whole-slide images (WSIs) from rectal biopsies. WSIs were tiled into 397,731 image patches (128x128 pixels), further partitioned into training (70%), validation (15%), and testing (15%) sets. Models were evaluated on tile- and patient-level diagnostic metrics and processing latency. Results. The U-Net achieved a tile-level sensitivity of 82.9%, showing no statistically significant difference compared to YOLO26 (79.1%; p = 0.097). However, YOLO26 demonstrated a statistically significant advantage in tile-level specificity (96.1% vs. 93.9%; p < 0.001) and reduced mean inference latency (7.64 ms vs. 11.57 ms/tile). At the patient level, both models achieved 100% diagnostic sensitivity. Despite low patient-level specificity (0.0% U-Net; 11.8% YOLO26), the tissue-level diagnostic burden of false positives was 6.00% for U-Net and 3.50% for YOLO26. Conclusion. The U-Net is preferred when nominal gains in sensitivity are prioritized, while the YOLO26 is an alternative that optimizes efficiency and false positive suppression. Both models serve as robust screening filters to augment the pathologist's workflow and should be selected based on workflow requirements. Prospective validation on larger, multi-centre datasets is required before clinical implementation.